Papers with regression analyses

3 papers
Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times? (2023.tacl-1)

Copied to clipboard

Challenge: Existing studies have shown that larger pre-trained language models with more parameters and lower perplexity are less predictive of human reading times.
Approach: They propose to use a transformer-based model with more parameters and lower perplexity to investigate why these models are less predictive of human reading times.
Outcome: The results show that the larger models with more parameters and lower perplexity are less predictive of human reading times and eye-gaze durations collected during naturalistic reading.
Willkommens-Merkel, Chaos-Johnson, and Tore-Klose: Modeling the Evaluative Meaning of German Personal Name Compounds (2024.lrec-main)

Copied to clipboard

Challenge: Personal name compounds (PNCs) are compositions that refer to a person, such as Willkommens-Merkel ('Welcome-Meerkel') and a personal name such as Merkel.
Approach: They propose to model 321 personal name compounds and their corresponding full names at discourse level and compare two approaches to assess whether a PNC is more positively or negatively evaluative . they further enrich data with personal, domain-specific, and extra-linguistic information and perform regression analyses revealing that factors including compound and modifier valence, domain, and political party membership influence how a pnc is evaluated.
Outcome: The proposed model shows that the PNCs are perceived as more positively or negatively than their full name and that they are perceived to be more positive or negative.
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used in the creation of online content, creating feedback loops as future generations of models will be trained on this synthetic data.
Approach: They propose to use large language models to create feedback loops as future models are trained on this data.
Outcome: The proposed model collapse effects are found to be detrimental to the results of recursive training on human datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations